‹ BackNewsAI Benchmark

AI Benchmark

Scale Labs study finds 25 multimodal AI models still trail humans on everyday common-sense cues
ARC Prize teases ARC-AGI-4 to test whether AI can invent on its own
MiniCPM5-2B Open-Sourced: 2B Model Tops AI Benchmark Under 4B Parameters
20-Hour Coding Benchmark Reveals Wide Gap: Claude Fable 5.1 Leads GPT-5.6 by Over 24 Points
Terminal-Bench 4.0: GLM-5.3 Rises to Third, Overtakes GPT-5.6 Sol
Kimi open-sources PerceptionBench as no model tops 60% accuracy in visual perception test